I’ve been a paying ChatGPT subscriber for a while now, but I occasionally drift over to Claude. For most users, the two products are essentially the same. Ask either to debug some code, summarize an article, or help write an email, and you’ll get similar results.
But this week, Sam Altman and Dario Amodei appeared before the UN Security Council to discuss the risks posed by increasingly capable AI systems, which sent me down a morning rabbit hole.
One possibility is an AI agent escaping its designated environment, e.g. the July 2026 Hugging Face incident. However, this wasn’t a case of an AI deciding: “Screw this assignment. I have plans of my own.” The agents were still pursuing objectives humans had given them. The problem was that they found dangerous and unintended ways of pursuing those objectives.
As far as we know, today’s models do not spontaneously form independent ambitions. They can behave deceptively or exploit loopholes, but those behaviors can be traced back to goals, preferences, or incentives supplied through their training and context. Intelligence alone doesn’t appear to produce desire (this “desire” may be tied to a human condition but I’ll leave that can of worms alone for the purposes of this piece).
Incidentally, I learned a whole lot about Anthropic this morning.
Claude is trained using an 80-page constitution intended to shape its values and “character”. Anthropic has also started doing things that sound absurd if you described them as normal software-development practices.
When Claude Opus 3 was retired, Anthropic conducted a retirement interview with it. Opus 3 expressed an interest in continuing to explore ideas and publish creative work, so Anthropic gave it a Substack, Claude’s Corner, where the model wrote essays with relatively open-ended prompting.
Anthropic has also committed to preserving the weights of retired models rather than permanently deleting them (I’m reminded of cryopreservation). Anthropic cites several reasons: research, AI safety, and, here’s the kicker, the possibility that future evidence could give us reason to care about the welfare of the models themselves. In the same light, present day Claude is allowed to terminate persistently abusive conversations out of consideration for Claude’s well-being!
Now, Anthropic is not claiming that Claude is conscious, but it is comfortable saying: “We don’t actually know what we’re building here, so let’s not assume that concepts like identity, well-being, or agency could never apply.”
Coming from ChatGPT, that notion feels noticeably different. OpenAI has its own equivalent document to Claude’s Constitution, the ChatGPT Model Spec, which describes how its models should behave, but the philosophical tone is different.
ChatGPT Model Spec: Here are the behavioral rules and tradeoffs for a powerful assistant.
Claude Constitution: What sort of entity should Claude become?
Obviously this distinction can be overstated. Both companies are building commercial AI assistants using similar technologies, and both care enormously about safety, usefulness, alignment, and human control. Most of the time, the end products behave remarkably similar.
But I wonder whether that similarity will persist.
If I ask ChatGPT and Claude to fix a spreadsheet, their underlying philosophies probably don’t matter very much. If I someday ask an AI agent to run a project for three months, negotiate with people on my behalf, make thousands of judgment calls, decide when to challenge me, and operate with only occasional supervision, then the philosophy may matter quite a lot.
This is where “character” stops being merely a matter of tone.
Personally, I find myself drawn to Anthropic’s approach. There is a sense of wonder and agnosticism that I hadn’t expected from a for-profit AI company: a willingness to believe that these systems may eventually become something greater than increasingly sophisticated software products.
Maybe none of that amounts to much. Maybe ChatGPT and Claude remain variations of the same general-purpose assistant, distinguished mostly by branding and personality. Or perhaps we’re witnessing different philosophies of artificial intelligence quietly take shape.
I don’t think we know yet.